Back

Current Research in Structural Biology

Elsevier BV

Preprints posted in the last 30 days, ranked by how well they match Current Research in Structural Biology's content profile, based on 12 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.

1
A High Throughput SPR-Based Array for Quantitative Profiling of Glycosaminoglycan Protein Interactions

Jowitt, T. A.; Birchenough, H. L.; Popplewell, J. F.; Dyer, D. P.; Day, A. J.

2026-07-04 biophysics 10.64898/2026.07.02.736113 medRxiv
Top 0.2%
0.6%
Show abstract

Glycosaminoglycans (GAGs) are linear, negatively charged, polysaccharides that mediate a wide variety of biologically critical interactions with proteins, underpinning growth factor signalling, extracellular matrix assembly and numerous disease processes. However, GAG-protein interactions remain under characterised, in part because of the lack of high-throughput tools to systematically profile binding across the GAG interactome. In this paper we present a novel Surface Plasmon Resonance-based array methodology utilising 16 commonly sourced GAG preparations (including chondroitin sulphate (CS), dermatan sulphate (DS), heparan sulphate, heparin, hyaluronan and keratan sulphate) allowing the specificity and affinity of GAG-binding proteins to be determined. As proof of principle, we have validated the array using four established GAG-binding proteins (antithrombin III, CD44, heavy chain 1 from inter--inhibitor and Slit2), generating data consistent with the known binding specificities and quantifying affinities for many of the interactions. The array also reveals previously unreported GAG interactions, including Slit2 binding to CS and DS, and CD44 binding to chondroitin sulphate E.

2
Structural Determinants of Catalytic Directionality in an AMP-Forming Acetyl-CoA Synthetase from Syntrophus aciditrophicus

Yaghoubi, S.; Dinh, D. M.; Thomas, L. M.; Wofford, N. Q.; McInerney, M. J.; Follmer, A. H.; Karr, E. A.

2026-07-07 biochemistry 10.64898/2026.07.06.736832 medRxiv
Top 0.3%
0.5%
Show abstract

Acetyl-coenzyme A (CoA) is a central metabolic intermediate that links carbon and energy metabolism across all domains of life. The conversion of acetate and acetyl-CoA is carried out by three enzyme pathways: acetate kinase/phosphotransacetylase, ADP-forming acetyl-CoA synthetase, and AMP-forming acetyl-CoA synthetase (Acs). Acs enzymes serve critical physiological roles across diverse organisms generally by catalyzing a reversible two-step reaction forming acetyl-CoA and AMP from acetate and ATP. Isolated from the wastewater reclamation facility in Norman, Oklahoma, Syntrophus aciditrophicus strain SB (Sa) relies on an AMP-forming acetyl-CoA synthetase (SaAcs1) that favors synthesizing acetate and ATP from acetyl-CoA and AMP, in contrast to all previously characterized Acs enzymes. The origin of this preference and the structural determinants of both the thioester-forming step and catalytic directionality remain poorly understood. Here, we report a 2.2 [A] crystal structure of full-length SaAcs1 in the adenylation conformation with acetyl-AMP bound in the active site. Structural comparison to the extensively characterized Acs enzymes from Salmonella enterica (SeAcs) and Cryptococcus neoformans (CnAcs) revealed a displaced CoA-binding loop in SaAcs1. Enzymatic assays confirmed that SaAcs1 preferentially catalyzes the ATP-forming reaction. Site-directed mutagenesis demonstrated that reversion of two residues, G196 and T197, at the beginning of the CoA-binding loop to the consensus sequence repositions the loop and shifts catalytic preference toward the AMP-forming direction. Together, these results establish the CoA-binding loop and G196 and T197 as the primary structural determinants of directional preference in SaAcs1.

3
Fast prediction of acidic amino acid sidechain conformations for cryo-EM modeling

Kolypetris, G.; Djurabekova, A.; Lasham, J.; Simsive, L.; Vonck, J.; Sharma, V.

2026-07-14 biophysics 10.64898/2026.07.12.738023 medRxiv
Top 0.3%
0.4%
Show abstract

Cryogenic-electron microscopy (cryo-EM) has revolutionized the field of protein structural biology. The structures of large membrane proteins are now routinely determined by cryo-EM to near atomic resolution. However, in the medium resolution range of cryo-EM maps (>[~]2 [A]), negatively charged sidechains of acidic residues are not well-resolved due to the negative electrostatic potential of the region. This may lead to incorrect sidechain models for residues like glutamic acid or aspartic acid that are central for proton transfer activity in various respiratory and photosynthetic enzymes. We previously proposed that the acidic residues with weak or non-existent cryo-EM density can be modeled to represent their low proton affinity conformations. Here, we tested this hypothesis on a larger data set of acidic amino acid residues in two high-resolution respiratory complex I structures. By using faster sidechain modeling and proton affinity prediction tools, we created a workflow that generates sidechain conformations of selected amino acid residues. We validated the sidechain conformation predictions by Q-score analysis and atomistic molecular dynamics simulations in different charged states. The proposed workflow provides a way to rapidly obtain sidechain conformations of acidic residues with weak cryo-EM densities and can be integrated into the existing cryo-EM modeling pipelines to speed up sidechain rotamer prediction.

4
Mitochondrial Lon protease couples substrate translocation to proteolytic activation

Schenck, N.; Ahrensback Roesgaard, M.; Abrahams, J. P.

2026-06-23 biochemistry 10.64898/2026.06.23.733973 medRxiv
Top 0.5%
0.3%
Show abstract

Human LonP1 is an ATP-dependent mitochondrial protease that degrades damaged or redundant proteins. Indiscriminate proteolysis by LonP1 is limited through tight coordination of substrate recognition, unfolding, translocation and catalytic cleavage, yet the role of ATP hydrolysis in these individual steps remains unclear. Here, we show that LonP1 binds substrates and cleaves peptide bonds without ATP hydrolysis, whereas degradation of folded proteins strictly depends on ATP-driven unfolding and translocation. Initial substrate binding opens a closed ADP-bound resting state, enabling nucleotide exchange and stimulating ATPase activity. The opening also increases accessibility of the proteolytic chamber, modestly enhancing peptidase activity. Maximal peptidase activity is observed in a transition-state mimic stabilised by ADP{middle dot}AlF, in which substrate is engaged within the translocation channel. Cryo-EM analysis reveals that in this state the proteolytic active sites are no longer occluded, linking ATP-driven substrate translocation to full proteolytic activation. Together, these findings reveal how LonP1 prevents indiscriminate proteolysis during substrate selection by ensuring that efficient proteolysis occurs only in substrate-translocating states. Model of the conformational landscape and functional cycle of LonP1Schematic overview of LonP1 states and their inter-conversion. State transitions are modulated by substrate, nucleotide occupancy, temperature, and inhibitors. Key distinguishing features include the presence or absence of the lateral gap, nucleotide state, substrate engagement within the A-tunnel, and the handedness of the ATPase (A) domains. Additional indicators include the compactness of the proteolytic (P) domain and the presence of substrate density within the N-terminal (N) domain or at the coiled-coil domain (CCD) as well as the position of a loop within the catalytic centre. The depicted cryo-EM structures represent a model of a continuous conformational landscape and correspond to the closest matching biological states and positions within the reaction cycle, but may also capture transient intermediates or conformations stabilised by experimental conditions. The shown atomic models correspond to the states highlighted in larger font (R-state: PDB 7NGL; P1-state: PDB 7NFY; P2-state: PDB 7NGC; closed LonP1-ADP-substrate: PDB 9CC1). O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=118 SRC="FIGDIR/small/733973v1_ufig1.gif" ALT="Figure 1"> View larger version (59K): org.highwire.dtl.DTLVardef@16e0491org.highwire.dtl.DTLVardef@1ee02b1org.highwire.dtl.DTLVardef@f2b47aorg.highwire.dtl.DTLVardef@26f6b2_HPS_FORMAT_FIGEXP M_FIG C_FIG

5
Structural and Biochemical Analysis of the CABIT1 Domain of THEMIS

Negron Teron, K. I.; Ortiz-Salazar, D.; Beyett, T. S.

2026-06-25 biochemistry 10.64898/2026.06.24.734275 medRxiv
Top 0.5%
0.3%
Show abstract

T cells are important components of the adaptive immune system and develop through a selection process regulated by signaling through the T-cell receptor (TCR). Thymocyte-Expressed Molecule Expressed in Selection (THEMIS) is a TCR-proximal protein that modulates the activity of Shp1 phosphatase to influence TCR signaling during development. THEMIS has been shown to both activate and inhibit Shp1, but the molecular mechanisms of these functions are poorly understood. THEMIS contains two rare Cysteine All-Beta In THEMIS (CABIT) domains, the N-terminal of which interacts with Shp1 and is likely responsible for modulation of its phosphatase activity. Herein, we report the first crystal structure of the THEMIS CABIT1 domain. While a portion of the CABIT1 domain is poorly resolved, it appears to share the same overall fold observed in our recent CABIT2 crystal structure and AlphaFold predictions. We show that phosphorylation of the CABIT1 domain by LCK is required for association with SHP1 and that phosphorylated CABIT1 can protect Shp1 from oxidation and inhibition by reactive oxygen species (ROS), which may serve as a mechanism by which THEMIS enhances Shp1 activity.

6
The structure of the lipid II flippase from monoderm bacteria

Li, Y. E.; Baron, G. F.; Clemons, W.

2026-06-23 biochemistry 10.64898/2026.06.21.733627 medRxiv
Top 0.5%
0.3%
Show abstract

Peptidoglycan biogenesis requires membrane flippases to translocate lipid-linked precursors across the cytoplasmic membrane for processing (1). This essential step is mediated by MurJ, the lipid II flippase conserved across all peptidoglycan-producing bacteria (2). While MurJ from diderm bacteria has been structurally resolved in multiple conformational states (3-6), its monoderm homolog remains uncharacterized. Monoderm MurJ homologs exhibit substantial sequence divergence yet retain the same lipid II flipping function (7) and are promising antibiotic targets. Here we report structures of Staphylococcus aureus MurJ (SaMurJ) captured in both outward- and inward-facing conformations. These structures show that SaMurJ adopts the conserved MOP family fold and undergoes conformational transitions consistent with an alternating-access mechanism. Our findings reveal conserved and divergent features of MurJ between diderm and monoderm bacteria that are critical for lipid II flipping and provide a structural framework for probing substrate recognition and specific inhibition. Significance StatementThe growing global threat of antibiotic resistance and the limited development of new antibacterial therapies underscore the urgent need to identify and mechanistically characterize new antibiotic targets and mechanisms. MurJ is an essential membrane transporter required for cell wall biosynthesis and represents an attractive but unexplored antibiotic target. Here we determine the structures of MurJ from a clinically critical monoderm pathogen Staphylococcus aureus in key conformational states during its transport cycle. This work advances our understanding of an essential step in bacterial cell wall synthesis, reveals key distinctions between monoderm and diderm MurJ, and defines structural features that can be exploited for antibiotic discovery.

7
AlphaFlex: Ensembles of the human proteome representing disordered regions

Liu, Z. H.; Zhang, O.; De Castro, S.; Sun, K.; Ghafouri, H.; Attafi, O. A.; Fawzi, N. L.; Tosatto, S. C. E.; Monzon, A. M.; Moses, A. M.; Head-Gordon, T.; Forman-Kay, J. D.

2026-06-23 biochemistry 10.1101/2025.11.24.690279 medRxiv
Top 0.6%
0.3%
Show abstract

More than two thirds of proteins in the human proteome are predicted to contain intrinsically disordered regions (IDRs), which lack stable folded structure. IDRs are critical for biological regulation and organization, as targets for post-translational modifications, and as mediators of biomolecular condensates. To address the pressing need for better structural models enabling functional insight, we developed AlphaFlex to model fully atomistic conformer ensembles for proteins predicted to have IDRs, modeled in the context of AlphaFold folded domains and an implicit bilayer for transmembrane proteins. The AlphaFlex resource provides conformational ensembles of human proteins from the AlphaFold database with identified IDRs in the Protein Ensemble Database that is mirrored in UniProt. This transformative resource of AlphaFlex ensembles provides physically and biologically relevant full-length models for IDR proteins, including scaffold proteins, those with IDR:folded-domain interactions, regulatory and condensate proteins requiring exposed binding elements, conditionally folding IDRs, and transmembrane proteins containing IDRs.

8
Characterization of the trimeric TOM complex by HS-AFM single-molecule analysis

Kobayashi, N.; Omura, S. N.; Kuzasa, K.; Imai, K.; Kawai, S.; Imai, H.; Amyot, R.; Umeda, K.; Nureki, O.; Endo, T.; Kodera, N.; Araiso, Y.

2026-07-01 biochemistry 10.64898/2026.07.01.735793 medRxiv
Top 0.6%
0.3%
Show abstract

The translocase of the outer mitochondrial membrane (TOM) complex is the main entry gate for mitochondrial proteins. Approximately 99 % of mitochondrial proteins are synthesized as precursor proteins (preproteins) in the cytosol and subsequently translocated into mitochondria through the TOM complex. The TOM complex exists in a dynamic equilibrium among multiple assembly states through spatial rearrangements of its subunits. The recent cryo-electron microscopy (cryo-EM) studies revealed near-atomic structures of the TOM core dimer, whereas previous biochemical studies indicated the TOM complex functions as a trimer in intact mitochondria. However, the relationship between the core dimer and the functional trimer remains unclear. In the present study, we analyzed the dynamics of the TOM complex using high-speed atomic force microscopy (HS-AFM) to investigate the assembly states and conformation transitions of the TOM complexes. We demonstrated that purified yeast TOM complexes predominantly adopt a trimeric organization but dynamically dissociate into dimeric and monomeric states during HS-AFM observation. The trimeric particles observed by HS-AFM exhibited spherical molecular shapes consistent with a trimeric structural model proposed from previous crosslinking analyses. In contrast, the dissociated dimeric particles closely resembled the dimensions of the TOM core-dimer structures determined by cryo-EM. Furthermore, HS-AFM analyses provided insight into the spatial arrangement of the Tom20 receptor, consistent with previous models of the trimeric TOM complex. These observations enabled characterization of the trimeric TOM complex in vitro and provide a foundation for future structural and functional analyses of TOM complex assembly.

9
Direct Binding of Cysteine-367 Thiolate to the Active Site of the -Hydrogenase from Clostridium beijerinckii in the O2-stable State

Duan, J.; Arrigoni, F.; Rutz, A.; Hofmann, E.; Greco, C.; Happe, T.

2026-07-13 biochemistry 10.64898/2026.07.11.737921 medRxiv
Top 0.7%
0.2%
Show abstract

[FeFe]-hydrogenases are very active biocatalysts for H2 conversion. However, their active site is vulnerable to irreversible degradation initiated by O2 binding at the catalytic iron ion (Fed) of the active center. CbA5H, the [FeFe]-hydrogenases from Clostridium beijerinckii exhibits stability towards oxygen (O2) due to its ability to reversibly enter an inactive state termed Hinact upon contact with O2. We previously proposed that the close distance of approximately 3.1 [A] between the thiol of a nearby cysteine (C367) and the Fed, based on a 2.9 [A] crystal structure of CbA5H in the Hinact state, enables their binding to each other. This binding therefore was suggested to shield the Fed from O2 damage. However, there is currently a lack of evidence to support this hypothesis. Furthermore, density functional theory (DFT) calculations based on a homologous model favored hydroxide as the binding ligand of the Fed over the thiol of C367. In this study, we present the crystal structure of CbA5H in the Hinact state at an improved resolution of 2.15 [A]. The structure reveals a direct binding between the thiol of C367 and the Fed with a distance of approximated 2.77 [A] which is well supported by our DFT calculations based on the new crystallographic data. It is noteworthy that the 2.77 [A] bond distance is strikingly long when compared with other iron-sulfur bonds. This finding may provide a crucial foundation for understanding the rapid reversibility of the Hinact state.

10
Structures of the human sodium-citrate cotransporter NaCT with and without substrates

Sauer, D. B.; Song, J.; Marden, J. J.; Wang, B.; Sowerby, K.; Sudar, J. C.; Rice, W. J.; Wang, D.-N.

2026-07-12 biophysics 10.64898/2026.07.08.737274 medRxiv
Top 0.7%
0.2%
Show abstract

The human sodium-citrate cotransporter NaCT imports various tri- and dicarboxylates into the cell as TCA cycle intermediates. This substrate uptake process is driven by an inward sodium gradient. The protein is a member of the Divalent Anion-Sodium Symporter (DASS) family. Whereas extensive biochemical and structural studies have been carried out for NaCT, how the substrate binding and translocation is coupled to the sodium gradient remains unclear. Here using single particle cryo-electron microscopy, we determined the structures of the human NaCT protein in three states: sodium-free, in the presence of sodium, and sodium- and substrate-bound. These structures suggest a simultaneous binding mechanism for sodium-substrate coupling, distinct from the sequential binding, conformational selection mechanism previously observed for the bacterial DASS protein VcINDY.

11
Computational design of a multi-epitope vaccine against M. tuberculosis

Buhari, A.; Okutu, P.; Oyeleke, U. A.; Sivakumar, A.; Hameed, S. A.

2026-07-15 bioinformatics 10.64898/2026.07.09.737463 medRxiv
Top 0.7%
0.2%
Show abstract

BackgroundTuberculosis remains a leading global infectious killer, with BCG offering inconsistent adult protection and rising drug-resistant strains demanding novel vaccine strategies. We report the first multi-epitope vaccine construct simultaneously targeting three previously unexplored Mycobacterium tuberculosis virulence proteins; EccB3, MycP, and polyketide synthase which collectively govern nutrient acquisition, ESX secretion integrity, and innate immune evasion. MethodsUsing a reverse vaccinology pipeline, B-cell, CTL, and HTL epitopes were predicted, filtered for allergenicity, toxicity, and IFN-{gamma} induction, then assembled into an 823-residue chimeric construct incorporating beta-defensin and PADRE adjuvants with AAY/GPGPG linkers, covering [~]90% global HLA diversity. The construct underwent AlphaFold structure prediction, 3DRefine refinement, disulfide engineering, PROCHECK/ProSA validation, ClusPro 2.0 docking against TLR1/TLR2, and C-IMMSIM immune simulation. ResultsThe construct (82.3 kDa, instability index 32.48) showed strong structural quality (94.7% favoured Ramachandran residues), stable TLR1/TLR2 binding (weighted energy: -1,371.0 kcal/mol), and robust in silico immune responses and durable memory cell formation following booster simulation. ConclusionThis computationally validated construct represents a promising multi-target TB vaccine candidate warranting experimental advancement.

12
Structural Organization of the Nvj3-Mdm1 Complex Reveals a Conserved Lipid-Compatible Contact Site Module

Aboumourad, M.; Hariri, H.

2026-07-03 bioinformatics 10.64898/2026.06.29.735323 medRxiv
Top 0.8%
0.2%
Show abstract

Membrane contact sites are organized by protein assemblies that physically couple organelles and coordinate lipid metabolism, yet the structural principles that enable lipid exchange across these junctions remain poorly defined. At the nuclear-vacuolar junction (NVJ) in budding yeast, the tethering protein Mdm1 and its binding partner Nvj3 form a complex that regulates lipid metabolic pathways, but the structural features underlying their interaction have not been resolved. Here, we use AlphaFold-based complex prediction and comparative structural analysis to define the organization of Nvj3-Mdm1 complex assembly. We identify a high-confidence heterodimer in which conserved PXA and PXC domains generate an extended tunnel spanning both proteins. Tunnel analysis predicts a core hydrophobic conduit traversing the Nvj3-Mdm1 interface, consistent with a lipid-compatible architecture. Evolutionary conservation is enriched at the Nvj3-Mdm1 interface. The predicted conduit shares geometric and physicochemical properties with bridge-like lipid transfer proteins, including Atg2, Fmp27, and Hob2, suggesting that heteromeric tether assemblies may contribute directly to inter-organelle lipid transfer. Cophylogenetic analysis reveals coordinated coevolution of Nvj3 and Mdm1 across Saccharomycetes. Together, these findings define Nvj3 as a structural partner of Mdm1 and support a conduit-based model of lipid transfer at the NVJ.

13
Rhizobial enzyme reveals pH-driven catalytic switching and ʟ-amino acid incorporation by ʟ,-transpeptidases

Rady, B. J.; Bahadur, R.; Evans, C. A.; Mesnage, S.

2026-07-01 biochemistry 10.64898/2026.06.30.735398 medRxiv
Top 1%
0.1%
Show abstract

Nearly all bacteria are surrounded by a mesh-like macromolecule called peptidoglycan that gives them their shape and helps them resist turgor pressure. To grow and maintain their peptidoglycan, bacteria produce a wide range of enzymes, including the relatively understudied ,[x1D05]-transpeptidase (LDT) family. LDTs can catalyse several different reactions and vary widely in copy number: some bacteria have none, whilst others have more than twenty. To better understand why some bacteria have so many LDTs, we examined 18 putative ones from Rhizobium johnstonii, a nitrogen-fixing, symbiotic bacterium. Heterologous expression revealed several highly active enzymes, one of which, LdtRj8, we further characterized in detail. In vitro assays showed that LdtRj8 was capable of ,[x1D05]-transpeptidation, carboxypeptidation, substitution, and endopeptidation, but that its preferred activity differed at different pHs. LdtRj8 particularly excelled at ,[x1D05]-substitution, utilizing all of the tested [x1D05]-amino acids, and, surprisingly, most of the -amino acids as well. LdtRj8's pH-modulated activity could help R. johnstonii respond to acidic conditions encountered throughout the rhizobium-legume symbiosis, and its -amino acid substitution activity, which we show to be a more general property of LDTs, may regulate ,[x1D05]-transpeptidation and explain the existence of isomeric muropeptides often reported in the literature.

14
SAPPTree: Identification of an S-Acylation Motif Drives a Novel S-Acylation Prediction Program

Guo, A. S.; Luong, V. S.; Petropavlovskiy, A. A.; Dang, A.; Doxey, A. C.; Sanders, S. S.; Martin, D. D. O.

2026-07-03 bioinformatics 10.64898/2026.06.30.735287 medRxiv
Top 1%
0.1%
Show abstract

S-acylation, the reversible addition of fatty acids to proteins, has emerged as an abundant post-translational modification that drives protein localization and function. With no known consensus sequence, current prediction programs rely on machine learning algorithms that use short peptide sequences and large proteomic datasets. However, current prediction programs often suggest incorrect sites of S-acylation, leading to wasted experimental time and effort following site-directed mutagenesis and low-throughput validation experiments. Using only experimentally confirmed sites of S-acylation, we sought to identify primary sequence, secondary structure, and tertiary structure features common amongst S-acylation sites to aid in developing more robust prediction tools. In doing so, we identified an S-acylation motif including a cysteine cluster flanked by a hydrophobic stretch, and a positively charged polybasic region found within a helical stretch. These features were combined with known or AlphaFold-predicted structures and additional features including residue depth and solvent accessibility into a random forest model to generate a new and more accurate S-acylation prediction program (SAPP), named SAPPTree. All the processed datasets and complete model training pipeline are available at https://github.com/neurdyphagy-lab/palm-prediction-model, while the webserver is available at http://martintools.sci.uwaterloo.ca/.

15
Evidence for lanthanide and PQQ dependent dehydrogenases in Eukarya

Robinson, C. M.; Martinez-Gomez, N. C.; West-Roberts, J. A.; Voutsinos, M. Y.; Banfield, J.

2026-07-14 bioinformatics 10.64898/2026.07.14.738520 medRxiv
Top 1%
0.1%
Show abstract

Lanthanides function as enzyme cofactors in bacteria, where they are widely distributed in pyrroloquinoline quinone-dependent 8-bladed beta-propeller dehydrogenases. No lanthanide-dependent enzymes, however, have been described outside prokaryotes. Here, we combined structural bioinformatics, phylogenetics, AlphaFold3 co-folding, coordination-sphere comparison, and quantum-mechanical cluster modeling to search for and rank putative lanthanide-coordinating 8-bladed beta-propeller enzymes in Eukarya. We identified candidate lanthanide-coordinating proteins in a diverse range of eukaryotes, predominantly plants and fungi, including species of clear industrial and agricultural relevance. A high-confidence subset matched validated bacterial Ln-binders based on both geometric similarity to canonical Ln-binding sites and on predicted Ln3+ versus Ca2+ selectivity. Our findings indicate that lanthanide biology likely extends beyond bacteria, with implications for plant, fungal, and broader eukaryotic metabolism, and warrant targeted biochemical investigation.

16
The Gompertz curve for estimating growth rates of Protein Data Bank and protein folds

Sato, K.; TOMII, K.

2026-06-26 bioinformatics 10.64898/2026.06.24.732253 medRxiv
Top 1%
0.1%
Show abstract

The Protein Data Bank (PDB) is an ever-growing, open-access repository of structural data of biological molecules. This international database has been instrumental in the development of artificial intelligence and deep learning models for protein structure prediction and design. The PDB growth is a crucially important factor influencing further development of these models. Therefore, after analyzing the growth trend in PDB depositions since the archive's launch, we found that it is well fitted by the Gompertz function, a growth curve used across various disciplines. Furthermore, we observed that the function captures the "discovery of novel folds", i.e., the cumulative number of distinct folds among protein domains that constitute most of the PDB. Consequently, based on the fitting results, we estimated the likely numbers of PDB entries and protein folds. These findings provide insights into deceleration of growth in recent years and enable us to assess anticipated trends.

17
Structural Bioinformatics of Four Human Aquaporins and Their Water-Soluble QTY Analogs

Zhang, S.; Xiao, E.

2026-06-30 bioinformatics 10.64898/2026.06.24.734367 medRxiv
Top 1%
0.1%
Show abstract

Human aquaporins (AQPs) are essential membrane channels, yet their inherent hydrophobicity complicates structural and functional studies. We present the systematic application of the QTY code to human AQPs, integrating it with AlphaFold 3 structure prediction to design and validate that four-representative human AQPs (AQP1, AQP3, AQP4, AQP7) can be converted into water-soluble analogs while maintaining their conformation. This approach features a novel platform for editing challenging membrane proteins. The QTY code was applied to the transmembrane regions of the selected four AQPs. Subsequently, the water-soluble QTY analogs of the four AQPs were predicted using AlphaFold 3. The predicted structures were superposed with CyroEM- or X-ray-determined native structures in PyMOL. Further analyses included root-mean-square deviation (RMSD) calculations, visualization of hydrophobic surface reduction, and inspection of conserved protein-ligand binding ability. After applying the QTY code, sequence changes between native AQPs and their QTY analogs was significant (42.86-48.80%). Nevertheless, their structures superposed well in analyses, with only slight deviations (RMSD < 0.6 [A]). In addition, the surface hydrophobicity of all QTY-edited AQPs was significantly reduced. Importantly, molecular contacts between the cholesterol ligand and protein were largely preserved for both native AQP1 and its QTY analog. Finally, all AlphaFold3-predicted structures for AQPs have high confidence values (pLDDT > 90; pTM ~0.83), supporting the reliability of the predicted structures. The findings demonstrate that membrane protein hydrophobicity can be edited and reduced without compromising fold integrity or functional architecture. Integration of the QTY code with AlphaFold 3 affords a high-throughput platform for designing water-soluble, structurally faithful analogs of challenging membrane proteins. Such a strategy can provide a potent platform for detergent-free biochemical studies and water-soluble analogs for therapeutic monoclonal antibody discoveries, thus advancing research of this pharmacologically important protein family.

18
Discovery and structural analysis of glycoside hydrolase family 176 α-1,2 glucosidase from Arthrobacter humicola A8F5

Yasukochi, R.; Suzuki, T.; Toraya, T.; Hino, K.; Mori, T.; Kashima, T.; Miyanaga, A.; Watanabe, H.; Fushinobu, S.

2026-07-03 biochemistry 10.64898/2026.07.01.735942 medRxiv
Top 1%
0.1%
Show abstract

Glycoside hydrolases (GHs) exhibit remarkable specificity dictated by the structural configuration of their target glycosidic linkages. While enzymes that process -1,4- and -1,6-linkages in starch or glycogen are well-characterized, those acting on less common bonds, such as -1,2-glucosidic linkages, remain largely underexplored. In this study, we report the discovery and structural elucidation of a novel -1,2-glucosidase from Arthrobacter humicola A8F5 (A8F5 glucosidase), representing a newly uncovered activity within the poorly characterized GH176 family. Biochemical characterizations revealed that A8F5 glucosidase exclusively cleaves -1,2-linkages via an anomer-inverting mechanism, with a distinct preference for short kojioligosaccharides. To circumvent crystallization obstacles caused by high loop flexibility and translational non-crystallographic symmetry, we engineered a loop-truncated variant. This strategy enabled the determination of high-resolution (up to 1.79 [A]) crystal structures of the enzyme in its ligand-free form and in complex with kojibiose, kojitriose, and selaginose. A8F5 glucosidase adopts a (/{beta})6-barrel fold characteristic of clan GH-G. Complementing the crystal structures with AlphaFold3 prediction demonstrated that two prominent active-site loops (loops 3 and 4) adopt a closed conformation that constricts the catalytic pocket, rendering the architecture suitable for short oligosaccharide recognition while restricting access to larger polymers. Furthermore, sequence similarity network analysis highlights vast, uncharacterized functional diversity within the GH176 family. These findings revealed that the GH176 enzyme recognizes and hydrolyses -1,2-glucosidic bonds through a structural framework distinct from that of the previously known clan GH-L GH65 kojibiose hydrolase, expanding the known functional landscape of this enzyme group toward rare -glucans.

19
The 329HHK331 motif is essential for Alzheimer's disease tau filament fold

Sato, Y.; Kawasaki, M.; Moriya, T.; Senda, M.; Masuda-Suzukake, M.; Ando, K.; Hisanaga, S.-i.; Hasegawa, M.; Senda, T.; Nonaka, T.

2026-07-03 neuroscience 10.64898/2026.06.29.735439 medRxiv
Top 1%
0.1%
Show abstract

Cryo-electron microscopy (cryo-EM) has revealed disease-specific tau filament folds, yet the local sequence elements that determine them remain poorly understood. Here we focused on the 329HHK331 motif near an inter-protofilament interface in Alzheimer's disease (AD)-type tau filaments, and analyzed recombinant dGAE filaments of wild-type (WT) and mutants in this motif. All mutants formed amyloid-like filaments in vitro, but their morphologies differed. In cultured cells, WT filaments efficiently seeded WT tau aggregation. Filaments with three-residue changes (deletion or Ala substitution) showed almost no seeding activity, whereas two-residue deletions retained partial activity. Cryo-EM showed that WT dGAE filaments form a quadruple helical filament of two protofilament dimers. Each dimer comprises two C-shaped protofilaments, centered on a 333GGG335-mediated inter-protofilament interaction and supported by flanking 329HHK331-336QVE338 contacts. Three-residue alterations abolished interactions with the QVE motif at the protofilament interface, thereby displacing 333GGG335 and forming non-C-shaped protofilament structures that are intrinsically poor templates for tau seeding. By contrast, two-residue deletions maintained the C-shaped protofilament structure because the remaining residue formed alternative inter-protofilament interactions. These findings suggest that the 329HHK331 region is a key determinant of the AD-like C-shaped protofilament fold and link this motif to tau seeding, providing insight into disease-specific tau filament formation.

20
A comprehensive analysis of calreticulin mutants reveals distinct biophysicochemical proprieties with a potential for refined targeted therapies

Kurt, O. N.; Civelek, E.; Ozturk, B.; Chachoua, I.

2026-06-24 bioinformatics 10.64898/2026.06.19.733337 medRxiv
Top 1%
0.1%
Show abstract

Calreticulin mutations in myeloproliferative neoplasms result in the replacement of the C-terminus acidic sequence with a positively charged tail that causes pathological activation of the thrombopoietin. The two canonical variants are Type-1 and Type-2. The remaining are mainly classified as Type-1 or Type-2 like based on the wild type sequence retained. Here, we performed in silico biophysicochemical analyses of 76 CALR exon 9 frameshift variants by their sequence and predicted biophysical properties, complemented by structural modeling of the mutant homodimers. Beyond confirming the Type-1 versus Type-2 distinction, we found that the Type 1-like variants form a continuum of charge architecture along which two reproducible subgroups can be identified, rather than sharply separated classes. This work refines the conventional mechanism-based classification into a charge-resolved framework and provides testable hypotheses linking novel-tail chemistry to receptor activation in CALR-mutant neoplasms and paves the way for improved targeted therapies based on individual mutants characteristics